Using Runtime Measured Workload Characteristics in Parallel Processor Scheduling

نویسندگان

  • Thu D. Nguyen
  • Raj Vaswani
  • John Zahorjan
چکیده

We address the design of practical scheduling policies for scalable, shared memory multiprocessors. In particular, we propose and evaluate experimentally processor allocation policies that use information about parallel job execution characteristics in making their decisions. In contrast to existing processor allocation work, which for the most part has relied on perfect information supplied before job execution, we require no a priori information, instead relying on runtime measurement of executing jobs. The experimental results we present validate the following observations: The use of runtime measurements of workload characteristics can signi cantly improve performance relative to disciplines that are oblivious to these characteristics, even given the inherent inaccuracies in the measurements and the overhead of the dynamic reallocations that approaches based on them require. Runtime measurements are su cient for the scheduler to achieve performance surprisingly close to that possible when perfect, a priori information is available. The primary performance loss, relative to the use of a priori information, is due to the transient poor decisions of the scheduler as it acquires information on the running applications, rather than to the overhead of measurement itself. Additionally, we consider both interactive environments, in which a response time directed scheduler is appropriate, and batch environments, in which maximizing useful instruction throughput is the primary goal. Our experiments are performed on a prototype implementation running on a 50-node KSR-2 shared memory multiprocessor. A nal novel aspect of our work is the use of both hand and compiler parallelized programs in our test workloads.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Adaptively Scheduling Parallel Loops in Distributed Shared-Memory Systems

Using runtime information of load distributions and processor affinity, we propose an adaptive scheduling algorithm and its variations from different control mechanisms. The proposed algorithm applies different degrees of aggressiveness to adjust loop scheduling granularities, aiming at improving the execution performance of parallel loops by making scheduling decisions that match the real work...

متن کامل

A Comparative Study of Real Workload Tracesand Synthetic Workload

Two basic approaches are taken when modeling workloads in simulation-based performance evaluation of parallel job scheduling algorithms: (1) a carefully reconstructed trace from a real supercomputer can provide a very realistic job stream, or (2) a exible synthetic model that attempts to capture the behavior of observed workloads can be devised. Both approaches require that accurate statistical...

متن کامل

A Comparative Study of Real Workload

Two basic approaches are taken when modeling workloads in simulation-based performance evaluation of parallel job scheduling algorithms: (1) a carefully reconstructed trace from a real supercomputer can provide a very realistic job stream, or (2) a exible synthetic model that attempts to capture the behavior of observed workloads can be devised. Both approaches require that accurate statistical...

متن کامل

A Comparison of Workload Traces from Two Production Parallel Machines

The analysis of workload traces from real production parallel machines can aid a wide variety of parallel processing research, providing a realistic basis for experimentation in the management of resources over an entire workload. We analyze a ve-month workload trace of an Intel Paragon machine supporting a production parallel workload at the San Diego Supercomputer Center (SDSC), comparing and...

متن کامل

Contention-Aware Scheduling of Parallel Code for Heterogeneous Systems

A typical consumer desktop computer has a multi-core CPU with at least two and up to eight processing elements over two processors, and a multi-core GPU with up to 512 processing elements. Both the CPU and the GPU are capable of running parallel code, yet it is not obvious when to utilize one processor or the other because of workload considerations and, as importantly, contention on each devic...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 1996